Abstract
Background: Remote patient monitoring in heart failure is synthesised as a single intervention, with pooled estimates drawn from up to 79 randomised trials in more than 31,000 patients. Individual landmark trials disagree sharply, and subgroup interaction tests reporting no difference between monitoring modalities are used to justify pooling across them. Methods: Systematic reviews and landmark randomised trials of remote monitoring in heart failure were eligible. The smallest ratio of risk ratios detectable by a subgroup interaction test was derived analytically at conventional significance and power across the plausible range of subgroup precision. Trial-level characteristics were compared descriptively. No new pooling was performed. Analyses were performed in Python 3. Results: Landmark trials span the full range of results within one intervention category: benefit in CHAMPION, IN-TIME and TIM-HF2, neutrality in TIM-HF and GUIDE-HF, and no benefit in Tele-HF and BEAT-HF. Pooled estimates nonetheless favour monitoring, with Cochrane relative risks for all cause mortality of 0.87 for structured telephone support and 0.80 for noninvasive telemonitoring, and a contemporary synthesis reporting no significant interaction by modality (all-cause mortality p = 0.80; heart failure hospitalisation p = 0.14). Derivation shows that an interaction test of this design cannot detect a difference between modalities smaller than approximately 1.5-fold at the most favourable plausible precision, and 2.7-fold at the least; the reported null therefore does not establish that a telephone call and an implanted pulmonary artery sensor are equivalent. TIM-HF and TIM-HF2, conducted by the same investigators in the same country with broadly similar technology, reached opposite conclusions, the principal difference being eligibility. Two of 59 poolable trials reported a formal rural-versus-urban subgroup analysis. Conclusions: Remote patient monitoring in heart failure is a category spanning interventions that differ by orders of magnitude in invasiveness, intensity and cost. The interaction tests cited in support of pooling across that category are underpowered by a wide margin, and their null results should not be read as evidence of equivalence. The comparison between TIM-HF and TIM-HF2 suggests that patient selection may matter as much as technology, which is a question the pooled literature is not designed to answer.
Keywords: Heart failure, Remote patient monitoring, Telemonitoring, Subgroup interaction, Statistical power, Patient selection